emnlp 2020
MaLA-500: Massive Language Adaptation of Large Language Models
Lin, Peiqin, Ji, Shaoxiong, Tiedemann, Jörg, Martins, André F. T., Schütze, Hinrich
Large language models have advanced the state of the art in natural language processing. However, their predominant design for English or a limited set of languages creates a substantial gap in their effectiveness for low-resource languages. To bridge this gap, we introduce MaLA-500, a novel large language model designed to cover an extensive range of 534 languages. To train MaLA-500, we employ vocabulary extension and continued pretraining on LLaMA 2 with Glot500-c. Our experiments on SIB-200 show that MaLA-500 achieves state-of-the-art in-context learning results. We release MaLA-500 at https://huggingface.co/MaLA-LM
Crossing the Conversational Chasm: A Primer on Natural Language Processing for Multilingual Task-Oriented Dialogue Systems
Razumovskaia, Evgeniia (Language Technology Lab, University of Cambridge, UK) | Glavas, Goran (Data and Web Science Group, University of Mannheim, Germany) | Majewska, Olga (Language Technology Lab, University of Cambridge, UK) | Ponti, Edoardo M. (Mila - Quebec AI Institute and McGill University, Canada) | Korhonen, Anna (University of Cambridge, UK) | Vulic, Ivan (Language Technology Lab, University of Cambridge, UK)
In task-oriented dialogue (ToD), a user holds a conversation with an artificial agent with the aim of completing a concrete task. Although this technology represents one of the central objectives of AI and has been the focus of ever more intense research and development efforts, it is currently limited to a few narrow domains (e.g., food ordering, ticket booking) and a handful of languages (e.g., English, Chinese). This work provides an extensive overview of existing methods and resources in multilingual ToD as an entry point to this exciting and emerging field. We find that the most critical factor preventing the creation of truly multilingual ToD systems is the lack of datasets in most languages for both training and evaluation. In fact, acquiring annotations or human feedback for each component of modular systems or for data-hungry end-to-end systems is expensive and tedious. Hence, state-of-the-art approaches to multilingual ToD mostly rely on (zero- or few-shot) cross-lingual transfer from resource-rich languages (almost exclusively English), either by means of (i) machine translation or (ii) multilingual representations. These approaches are currently viable only for typologically similar languages and languages with parallel / monolingual corpora available. On the other hand, their effectiveness beyond these boundaries is doubtful or hard to assess due to the lack of linguistically diverse benchmarks (especially for natural language generation and end-to-end evaluation). To overcome this limitation, we draw parallels between components of the ToD pipeline and other NLP tasks, which can inspire solutions for learning in low-resource scenarios. Finally, we list additional challenges that multilinguality poses for related areas (such as speech, fluency in generated text, and human-centred evaluation), and indicate future directions that hold promise to further expand language coverage and dialogue capabilities of current ToD systems.
Deep Transfer Learning & Beyond: Transformer Language Models in Information Systems Research
Gruetzemacher, Ross, Paradice, David
AI is widely thought to be poised to transform business, yet current perceptions of the scope of this transformation may be myopic. Recent progress in natural language processing involving transformer language models (TLMs) offers a potential avenue for AI-driven business and societal transformation that is beyond the scope of what most currently foresee. We review this recent progress as well as recent literature utilizing text mining in top IS journals to develop an outline for how future IS research can benefit from these new techniques. Our review of existing IS literature reveals that suboptimal text mining techniques are prevalent and that the more advanced TLMs could be applied to enhance and increase IS research involving text data, and to enable new IS research topics, thus creating more value for the research community. This is possible because these techniques make it easier to develop very powerful custom systems and their performance is superior to existing methods for a wide range of tasks and applications. Further, multilingual language models make possible higher quality text analytics for research in multiple languages. We also identify new avenues for IS research, like language user interfaces, that may offer even greater potential for future IS research.
The Clinical Applications of NLP: Workshop at EMNLP 2020
This paper focuses on text summarization in the context of medical dialogue. The idea is that when a patient has a conversation with a doctor, you want to be able to automatically summarize what transpired in the conversation. Doctor: "What have you been experiencing today?" Doctor: "How severe is the pain on a scale from 1 to 10" Then you want to be able to produce a transcript-like record of what happened during the visit. In this paper, the authors propose a model that can summarize these dialogues by filling in words using local structures.
Truth-Seeking in the Post-Truth Era: Tutorial at EMNLP 2020
The obvious method is to cross-reference a claim with existing facts to verify whether that claim is true. This is accurate and highly explainable, but it also comes at some significant costs. Not only does it assume that the claim is checkable (e.g., what if a world leader decided to say "the entire universe was created last Thursday"), but it also requires a huge database of evidence because the AI needs sufficient evidence on so many different areas in order to be able to verify an acceptable proportion of claims. The alternate method is to use the context of a claim to try to make an assumption on whether or not it's true. For example, if the claim was made by The Onion, we could probably assume it's not true.
NLPBT 2020 - Call for Papers
Humans interact with each other through several means (e.g., voice, gestures, written text, facial expressions, etc.) and a natural human-machine interaction system should preserve the same modality. However, traditional Natural Language Processing (NLP) focuses on analyzing textual input to solve language understanding and reasoning tasks, and other modalities are only partially targeted. This workshop aims to be a forum for both academia and industry researchers where new and unfinished research in the area of Multi/Cross-Modal NLP can be discussed. In particular, the focus of this workshop are (i) studying how to bridge the gap between NLP on spoken and written language and (ii) exploring how NLU models can be empowered by jointly analyzing multiple input sources, including language (spoken or written), vision (gestures and expressions) and acoustic (paralingustic) modalities. All deadlines must be considered at 11.59pm GMT-12 (anywhere on Earth).